Senior DevOps Engineer
Company: Kotak Securities Limited
Location: Mumbai
Experience: 5-9 Years
Role Overview
We are looking for a Senior DevOps Engineer to help build and operate the infrastructure supporting Kotak Securities' core technology rebuild. The role spans cloud infrastructure automation, Kubernetes, CI/CD, observability, release engineering, and production reliability.
Candidates should have strong depth in either Cloud and Infrastructure Automation or Observability and Release Engineering, with working knowledge of the other.
Key Responsibilities
- Build and operate AWS infrastructure including networking, load balancing, DNS, IAM, and private connectivity.
- Develop and maintain Terraform modules and provisioning pipelines.
- Manage Kubernetes clusters, upgrades, autoscaling, resource allocation, ingress, and networking.
- Implement secrets, certificates, workload identity, and credential rotation.
- Support datacenter migration, hybrid connectivity, capacity planning, and cutover readiness.
- Build and maintain CI/CD pipelines with canary, blue-green, and feature-flagged deployments.
- Operate GitOps workflows, environment promotion, rollback, and drift reconciliation.
- Implement distributed tracing, metrics, logging, synthetic monitoring, and OpenTelemetry standards.
- Improve alerting through SLOs, burn-rate monitoring, and noise reduction.
- Build incident management tooling and improve postmortem and escalation processes.
Required Skills
- 5+ years in DevOps, SRE, Platform, or Infrastructure Engineering.
- Strong production experience with Kubernetes.
- Hands-on experience with Terraform or similar Infrastructure as Code tools.
- Strong Linux, networking, TCP, DNS, filesystem, and performance troubleshooting knowledge.
- Experience with Python, Go, or Bash for automation.
- Strong understanding of containers, cgroups, namespaces, and runtime behaviour.
- Strong expertise in either cloud infrastructure automation or observability and release engineering.
- Good written communication for runbooks, design documents, and postmortems.
Preferred Skills
- Istio, Linkerd, Cilium, BGP, routing, and firewall policies.
- Postgres, Redis, Kafka, or MongoDB operations.
- Dynatrace, Datadog, New Relic, Prometheus, Grafana, Loki, Mimir, or OpenTelemetry.
- Chaos engineering, ServiceNow integration, load testing, and performance engineering.
- Experience with low-latency infrastructure, market data systems, kernel tuning, or exchange connectivity.
nterview Process
The process includes system design, a practical troubleshooting exercise, and discussions around production incidents and reliability improvements. Candidates should also share a GitHub profile, portfolio, or other examples of code they have written.